Papers by Aashish Anantha Ramakrishnan

4 papers
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)

Copied to clipboard

Challenge: Figure captions are crucial for helping readers understand and remember a figure’s key message.
Approach: They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure .
Outcome: The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles.
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis (2026.findings-acl)

Copied to clipboard

Challenge: Text-to-image (T2I) models are trained on literal, object-centric prompts designed to reflect the visible contents of an image.
Approach: They propose a method to extract key subjects and enhance their representation at embedding-level using Large Language Models.
Outcome: The proposed model significantly improves image-caption consistency and human preference alignment.
From Intentions to Techniques: A Comprehensive Taxonomy and Challenges in Text Watermarking for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are rapidly growing and allowing textual content to be protected against unauthorized use.
Approach: They present a unified overview of different perspectives behind designing watermarking techniques through a comprehensive survey of the research literature.
Outcome: The proposed methods are based on the evaluation datasets used and watermarking addition and removal methods to construct a taxonomy.
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships? (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on assessing factual and logical correctness in downstream tasks with limited emphasis on evaluating MLLMs’ ability to interpret pragmatic cues and intermodal relationships.
Approach: They propose to use Coherence Relations to assess MLLMs' ability to perform multimodal discourse analysis using different prompting strategies.
Outcome: The proposed model fails to match the performance of simple classifier-based benchmarks on 10+ MLLMs using different prompting strategies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations